fix(replay-guard): block targeted base-work reverts and test regressions regardless of size - #533
Conversation
…ons regardless of size google-labs-jules[bot] pushed stale pre-merge workspace snapshots onto open PR branches org-wide. The guard's bulk signature (>=5 removed files, >=500 deleted lines, 4:1 ratio) caught the bulk replays (seedream#92, html4tree#131) but missed appguardrail#297: a small targeted revert that dropped develop's accessibility wrapper and deleted its regression tests, staying under every bulk threshold. Add two size-independent block signals on the post-merge-anchor range: - Unmerged base work: any path changed since the merge anchor whose content is byte-identical to the pre-merge first parent (set difference of two name-only diffs). This is exactly "the push reverted what the base merge brought in", however small. - Test regression without replacement: post-merge commits delete or net-shrink test files while adding no new test file anywhere in the push. Renames/refactors that add a replacement test still pass. Failure reports now list every matching reason and name the offending paths. Script-only change; the opencode-review.yml invocation is untouched. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
OpenCode Review Overview
Changed-File Evidence Mapflowchart LR
PR["PR changed files"] --> Evidence["OpenCode bounded evidence"]
Evidence --> S1["CI script: pr_head_replay_guard.py"]
S1 --> I1["review and security gate shell path"]
I1 --> R1["Review risk: CI script: pr_head_replay_guard.py"]
R1 --> V1["bash -n plus Strix self-test"]
Evidence --> S2["Test: test_pr_head_replay_guard.py"]
S2 --> I2["regression suite"]
I2 --> R2["Review risk: Test: test_pr_head_replay_guard.py"]
R2 --> V2["targeted test run"]
|
There was a problem hiding this comment.
Pull request overview
Approval sufficiency: bounded evidence supplied affirmative approval evidence for changed files, coverage/docstring posture, risk surfaces, and current-head verification; approval is not based merely on the absence of known blockers.
Verification posture: CodeGraph evidence was initialized and bounded current-head evidence reviewed for changed-file evidence including scripts/ci/pr_head_replay_guard.py, tests/test_pr_head_replay_guard.py.
Linter/static: workflow/static review evidence is bounded by the current-head GitHub Checks gate and changed-file evidence.
TDD/regression: coverage execution evidence and focused changed hunks were reviewed from bounded-review-evidence.md.
Coverage: coverage execution evidence reports supported repository test suites passed.
Docstring coverage: coverage execution evidence reports configured repository docstring gates passed or docstring coverage was advisory.
DAG: CodeGraph/source-backed behavior map connects scripts/ci/pr_head_replay_guard.py to the affected review, runtime, or workflow path and required checks.
PoC/execution: coverage-evidence job executed on the current head and reported PASS.
DDD/domain: workflow and repository-governance invariants were reviewed against changed files in bounded evidence.
CDD/context: CodeGraph evidence, changed-file history, and focused hunks were reviewed from bounded-review-evidence.md.
Similar issues: changed-file history evidence was reviewed for comparable local precedents.
Claim/concept check: bounded evidence, repository source, current-head workflow evidence, and, where numeric, scientific, statistical, or literature-backed claims are affected, original-paper/formula evidence and parameter-recovery expectations were used for claims.
Standards search: standards and external-source checks are delegated to configured OpenCode web_search/Context7/DeepWiki sources when applicable; no evidence-backed standards blocker is present in bounded evidence.
Compatibility/convention: changed workflow/script conventions, object naming, and reserved-word safety for schema/API/config/code surfaces were checked in bounded evidence.
Breaking-change/backcompat: deployment evidence and changed-file history were checked for backward-compatibility risk.
Performance: changed surfaces were checked for performance risk in bounded evidence.
Developer experience: changed automation, review, test, setup, and maintenance surfaces were checked for helpful or obstructive DX impact in bounded evidence.
User experience: connected user, operator, API, CLI, documentation, review-comment, status-check, rendering, and workflow-reader behavior was checked for contradictions against code, docs, and tests in bounded evidence.
Visual/DOM: Playwright visual, DOM locator, ARIA snapshot, console, and responsive evidence were checked when a web UI surface was present; for non-web surfaces, API/CLI/log/docs/workflow interaction evidence was reviewed instead.
Accessibility/i18n: accessibility, localization, and human-readable text surfaces were checked where UI, CLI, API message, docs, logs, or review text changed.
Supply-chain/license: dependency, package, model, container, and external-tool changes were checked in bounded evidence.
Packaging: package, build, test, lint, and security contracts were checked in bounded evidence.
Security/privacy: workflow-token, review-gate, and repository-automation security/privacy boundaries were checked in bounded evidence.
Findings
No blocking findings.
Adversarial validation
{"status":"passed","probes":[{"path":"scripts/ci/pr_head_replay_guard.py","line":1,"hypothesis":"New logic fails to catch targeted reverts of base work.","attack_or_counterexample":"Simulated targeted revert of base work in a test case.","evidence":"Added test case in `tests/test_pr_head_replay_guard.py` verifies detection of targeted reverts.","outcome":"falsified"},{"path":"scripts/ci/pr_head_replay_guard.py","line":1,"hypothesis":"New logic fails to catch test regressions without replacement.","attack_or_counterexample":"Simulated deletion of test files without adding new ones.","evidence":"Added test case in `tests/test_pr_head_replay_guard.py` verifies detection of test regressions.","outcome":"falsified"}],"residual_risk":"Low; edge cases are well-covered by tests."}Evidence
- Result: APPROVE
- Reason: PR enhances replay guard to block targeted base-work reverts and test regressions, with comprehensive test coverage.
- Scope:
central OpenCode/Strix review-process - Changed files:
2 - Head SHA:
e534faa3457be81a113fabce258d88d4bcbdc310 - Workflow run: 29246845763
- Workflow attempt: 1
This approval path is limited to ContextualWisdomLab/.github central review-process self-repair.
Motivation
google-labs-jules[bot] has been pushing stale pre-merge workspace snapshots onto open PR branches org-wide. Confirmed damage:
The guard caught the bulk replays but had no size-independent signal for targeted reverts.
What this adds
Script-only hardening of
scripts/ci/pr_head_replay_guard.py(no workflow changes; the existing--repo-root/--base-sha/--head-shainvocation is unchanged). Two new size-independent block signals, evaluated on the post-merge-anchor range:git diff --name-onlyruns (anchor..headminusanchor^1..head). Catches the appguardrail#297 targeted revert regardless of diff size.tests/,test_*,*_test,*.spec.*,*.test.*) while adding no new test file anywhere in the push. Legitimate refactors that rename/replace tests add a test file and pass; a rename shows up as D+A, so it is not flagged.The exact-ancestor-tree replay check and the conservative bulk signature are unchanged. Every failure now reports all matching reasons and names the offending paths (capped at 10 with an overflow count), so the log states exactly which deleted tests / unmerged paths triggered the block.
False-positive posture
Evidence
python -m pytest -q --cov --cov-report=term-missing— full suite green,scripts/cicoverage 100% (fail_under=100 gate).python -m interrogate -c pyproject.toml scripts— PASSED (100.0% docstrings).🤖 Generated with Claude Code